Mixed-Methods Pipeline

Research Methodology

We employed a robust sequential mixed-methods design, moving from theoretical scoping and large-scale data scraping to educational interventions and depth analysis.

Literature Review

Foundation: Our theoretical framework identified three core dimensions of exclusion that guided the entire project design.

Socio-economic

Gender, Race, Finance

AI Literacy

Self-Efficacy Gaps

Structural

Policy & Funding

Read Full Literature Review

Data Scraping & NLP Pipeline

To empirically map the structural landscape, we developed a custom Natural Language Processing (NLP) pipeline.

  • Dataset: 2,643 live PhD listings scraped from FindAPhD.com.
  • N-Gram Analysis: Identified recurring linguistic patterns (Bigrams/Trigrams) in funding criteria.
  • Critical Insight: Revealed the "False Hope" barrier—where projects advertise as "Worldwide" but lack visa sponsorship funding.

Training Intervention

The "Getting Started in AI" course, co-designed with the Jean Golding Institute (JGI), served as a live laboratory to observe shifts in student self-efficacy.

Session 1: Foundations

2.5 Hrs
  • Demystifying the "Black Box"
  • Supervised vs. Unsupervised
  • LLMs & Transformers

Session 2: Hands-on Python

2.0 Hrs
  • Pandas Data Exploration
  • Scikit-learn Classifiers
  • Intro to PyTorch

Quantitative Surveys

A longitudinal Pre- and Post-test design tracked quantitative shifts in student attitudes before and after training.

Instruments Used

  • AILIT-S Framework: Adapted from Hornberger et al. (2025) to measure perceived technical competence.
  • Likert Scales: 5-point agreement scales (e.g., "I feel confident explaining AI concepts").

Qualitative Interviews

Following the surveys, we conducted semi-structured interviews (45-60 mins) via Microsoft Teams.

These sessions allowed us to probe the "why" behind the quantitative data, exploring personal narratives regarding career expectations and institutional trust that numbers alone could not reveal.

Data Analysis Strategy

The final phase involved synthesizing collected data streams to triangulate findings.

  • Qualitative (RTA): We employed Reflexive Thematic Analysis (Braun & Clarke). An abductive approach was used to base codes in extant literature while remaining open to novel student-driven themes.
  • Quantitative: Descriptive statistics and paired t-tests identified significant shifts in confidence pre/post training.

Ethical Compliance & Incentives

Approved by the Faculty of Science and Engineering REC (Ref: 2026-29619). We prioritised fair compensation for student time.

£10 Voucher (Survey) £20 Voucher (Interview) Anonymisation: 31/03/2026